Papers with automated evaluation methods

9 papers
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on temporal knowledge in text-to-image models have not explored how temporal phenomena are handled in text models.
Approach: They propose a data set to holistically evaluate temporal knowledge in image generation using 7.9k prompts and more than 600 reference images.
Outcome: The proposed model evaluates temporal knowledge in image generation using 7.9k prompts and more than 600 reference images.
ConQRet: A New Benchmark for Fine-Grained Automatic Evaluation of Retrieval Augmented Computational Argumentation (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for evaluating RAArg are costly and lack long, complex arguments and real-world evidence.
Approach: They propose to use multiple fine-grained LLM judges to evaluate RAArg using a new benchmark that features long and complex human-authored arguments on debated topics.
Outcome: The proposed methods provide better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing.
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing (2026.eacl-long)

Copied to clipboard

Challenge: a single prompt can inspire countless valid stories, making objective verification impossible.
Approach: They propose a large-scale benchmark for creative writing evaluation using a reddit corpus and a 2,480-pair test set.
Outcome: The proposed model outperforms existing OTS judges and generative reward models in the evaluation of creative writing.
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations (2023.acl-long)

Copied to clipboard

Challenge: Prior work has shown that models may exploit shortcuts that are difficult to detect using standard n-gram similarity metrics such as ROUGE.
Approach: They propose to use human-assessed summary quality facets and pairwise preferences to improve MDS evaluation methods.
Outcome: The proposed methods improve the quality of literature review summarization models . they use human-assessed summary quality facets and pairwise preferences .
Evaluating Saliency Explanations in NLP by Crowdsourcing (2024.lrec-main)

Copied to clipboard

Challenge: a crowdsourced method to evaluate saliency methods in NLP is proposed . saliencies are difficult for humans to understand, and can cause psychological harm .
Approach: They propose a method to evaluate saliency methods in NLP by crowdsourcing . they recruited 800 crowd workers and empirically evaluated seven salience methods .
Outcome: The proposed method evaluates saliency methods on two datasets using crowdsourced data . it shows that the results are comparable to existing methods on NLP and CV fields .
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research (2025.acl-long)

Copied to clipboard

Challenge: a benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research is available online.
Approach: They propose to use a benchmark to evaluate LLMs' ability to design ablation studies . they investigate whether current automated evaluation methods are not reliable .
Outcome: The benchmark compared leading LLMs with human experts on generating detailed ablation study designs . the results show that current evaluation methods are not reliable for the task .
Revisiting Automated Evaluation for Long-form Table Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing automated metrics for long-form table question answering (LFTQA) are poorly correlated with human judgments and fail to distinguish between factually accurate responses and those that are factual incorrect.
Approach: They propose to use a meta-evaluation dataset to assess the effectiveness of LLM-based LFTQA systems.
Outcome: The proposed meta-evaluation dataset includes 2,988 human-annotated examples.
Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for enhancing Large Language Models (LLMs) struggle with novelty and Reinforcement Learning from human feedback (RLHF) is costly.
Approach: They propose to use a Reward Model (RM) and a principle-guided LLM-as-a-Judge to enhance creative output over baselines.
Outcome: The proposed approach significantly enhances creative output over baselines, but the principle-guided LLM-as-a-Judge yields superior generation quality.
Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus (2023.emnlp-main)

Copied to clipboard

Challenge: Societal gender asymmetries and inequalities are perpetuated through language . MT often defaults to masculine representations by making undue binary gender assumptions .
Approach: They propose a benchmark and automated evaluation methods to assess gender-neutral translation from English to Italian.
Outcome: The proposed method is based on a survey on gender-neutral translation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations